Papers with speech dataset

2 papers
Crowdsourcing Speech Data for Low-Resource Languages from Low-Income Workers (2020.lrec-1)

Copied to clipboard

Challenge: Existing platforms collect labelled speech data from urban speakers whose dialects are often very different from low-income users.
Approach: They propose to collect labelled speech data directly from low-income workers . they collect 109 hours of data from 36 participants in the Marathi language .
Outcome: The proposed approach can provide valuable supplemental earning opportunities to low-income rural and urban workers.
TV-AfD: An Imperative-Annotated Corpus from The Big Bang Theory and Wikipedia’s Articles for Deletion Discussions (2020.lrec-1)

Copied to clipboard

Challenge: Detecting imperatives in oral and written communication is difficult when the user doesn't use the expected forms.
Approach: They created an imperative corpus with dialogues from The Big Bang Theory and Wikipedia comments from Wikipedia . they manually annotated imperatives and used a syntax-based classifier to extract 10,624 statements that may be imperative.
Outcome: The proposed model performs better in the written data compared to speech data, but has a low precision and recall for speech data.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations